Papers with base models
Copied to clipboard
| Challenge: | Neural models generate the most common and generic responses all the time . Empirical results show that our method can significantly improve the diversity of responses generated by sequence-to-sequence models. |
| Approach: | They propose an iterative training process and ensemble method based on boosting to improve the diversity of responses generated by neural models. |
| Outcome: | Empirical results show that the proposed method significantly improves diversity and relevance of responses generated by all models. |
Copied to clipboard
| Challenge: | generative large language models have become crucial for modern NLP research and applications across multiple languages. |
| Approach: | They introduce the GigaChat family of Russian LLMs, available in various sizes . they evaluate their performance on Russian and English benchmarks and compare them with multilingual analogs . |
| Outcome: | The proposed model family is available in various sizes and is tested on Russian and English benchmarks. |
Copied to clipboard
| Challenge: | Recent studies have focused on zero-shot cross-lingual transfer of pretrained languages. |
| Approach: | They propose to use few-shot cross-lingual transfer to improve zero-shot performance of multilingual pretrained language models. |
| Outcome: | The proposed model can be scaled to high-quality samples and improves on zero-shot performance. |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) prompting is the dominant strategy for eliciting step-by-step reasoning in large language models, but its effect on code generation is poorly understood. |
| Approach: | They develop a chain-of-thought (CoT) prompting router that selects among 12 prompt styles via a single 84 ms forward pass. |
| Outcome: | The proposed model outperforms CoT in small models with a 84 ms forward pass. |
Copied to clipboard
| Challenge: | Existing methods to protect PII from training on small corpora are difficult to implement in real-world applications. |
| Approach: | They propose an entity-based framework that synthesizes encrypted training data to protect PII. |
| Outcome: | The proposed framework outperforms base models and ensures PII security on limited-scale datasets while exhibiting a modest performance gap compared to models trained on unencrypted synthetic data. |
Copied to clipboard
| Challenge: | Existing instruction-tuned open-source LLMs have only been instruction- tuned for English and a few popular languages, thus hindering their accessibility to many other languages in the world. |
| Approach: | They propose a framework that uses supervised fine-tuning and reinforcement learning from human feedback to improve the accessibility of large language models. |
| Outcome: | The proposed framework enables the evaluation of generative LLMs in multiple languages. |
Copied to clipboard
| Challenge: | Large language models produce non-existing facts when faced with questions outside their parametric knowledge, which undermines their reliability. |
| Approach: | They propose a method that separates the learning of answer prediction and confidence estimation during fine-tuning on instruction data. |
| Outcome: | Experiments on multiple models and different model sizes show that the proposed method outperforms baselines by up to 25% in average precision. |
Copied to clipboard
| Challenge: | Existing research is limited by general or niche datasets that lack sufficient scale for training dialogue systems. |
| Approach: | They propose a synthetic dialogue generation framework that uses Large Language Models and Chain of Thought reasoning to generate dynamic, domain-specific dialogues with simulated personas and diverse conversational features. |
| Outcome: | The proposed framework outperforms existing frameworks on dialogue summarization and quality increases as the size of the LLM increases from 3B to 8B. |
Copied to clipboard
| Challenge: | Existing hierarchical text classification methods make local decisions regarding labels or ignore hierarchy information during inference. |
| Approach: | They propose to learn a Label Assignment Policy via deep reinforcement learning to determine where to place an object and when to stop the assignment process. |
| Outcome: | The proposed method outperforms state-of-the-art methods on five datasets and four base models and achieves an average improvement of 33.4% over flat classifiers. |
Copied to clipboard
| Challenge: | Most state-of-the-art large language models (LLMs) are trained mainly on English data, limiting their effectiveness on non-English, especially low-resource, languages. |
| Approach: | They train language adapters for 13 languages and evaluate their effectiveness on downstream tasks using either task adapters or in-context learning. |
| Outcome: | The proposed language adapters improve performance for languages not seen during pretraining, but provide negligible benefit for seen languages. |
Copied to clipboard
| Challenge: | Domain- and customer-specific requirements complicate the problem of NL2SQL customization. |
| Approach: | They propose a distilled customization framework tailored for NL2SQL tasks. |
| Outcome: | The proposed framework outperforms teacher models on three benchmarks and achieves an average improvement of 36% in execution accuracy. |
Copied to clipboard
| Challenge: | Recent advances in language model reasoning require computationally intensive reinforcement learning and massive datasets. |
| Approach: | They propose a framework that combines Direct Preference Optimization and Supervised Fine-Tuning with selective guidance from larger models and iteratively refining solutions through a "reflect, rewrite, repeat" cycle. |
| Outcome: | The proposed framework shows significant performance improvements across arithmetic, symbolic and cognitive reasoning benchmarks. |
Copied to clipboard
| Challenge: | a challenge for aspect term extraction is to extract phrase-level aspect terms . a constituency lattice structure is constructed using the span annotations of constituents of a sentence . |
| Approach: | They propose to incorporate the span annotations of constituents of a sentence to leverage syntactic information in neural network models. |
| Outcome: | The proposed model outperforms existing models on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing approaches to integrating commonsense knowledge into large language models are implicit and explicit. |
| Approach: | They analyze the effects of model size and methods of injecting knowledge into TellMeWhy datasets to determine what aspects of commonsense knowledge are available in large language models. |
| Outcome: | The largest models yield substantial improvements over base models, but the amount of improvement decreases with larger model size. |
Copied to clipboard
| Challenge: | a finetuned model may be better base models than the vanilla pretrained model . this scheme, often referred to as intertraining, is the focus of the present work . |
| Approach: | They propose a scheme to analyze the potential intertraining gain independently for the target dataset and for a base model being considered as a starting point. |
| Outcome: | The proposed model is strong even if training data was not aligned with target dataset. |
Copied to clipboard
| Challenge: | Existing studies on instruction following focus on simple instructions and short responses . however, there are challenges associated with collecting preference judgments on long-form texts . |
| Approach: | They propose an instruction-following alignment method that uses dispreferred instructions to obtain negative feedback from dispvoted instructions. |
| Outcome: | The proposed model generates significantly longer texts than base models without significant quality degradation. |
Copied to clipboard
| Challenge: | Using auxiliary functions to implement functions is important for instruction-tuned models because it reduces the implementation difficulty of a target function compared to implementing them from scratch. |
| Approach: | They propose several ways to provide auxiliary functions to the models by adding them to the query or providing a response prefix to incorporate the ability to utilize auxiliary function with the instruction following capability. |
| Outcome: | The proposed models outperform the recent powerful language models, gpt-4o, in the code generation task. |
Copied to clipboard
| Challenge: | Autoregressive text style transfer models often ignore part of the source sentence and generate some irrelevant words with strong styles. |
| Approach: | They propose a non-autoregressive generator for unsupervised text style transfer which explicitly models word alignments to suppress irrelevant words. |
| Outcome: | The proposed generator significantly improves performance and provides explainable word alignments. |
Copied to clipboard
| Challenge: | Propaganda detection in social media is challenging due to noisy, short texts and low annotation agreements. |
| Approach: | They propose a new intent-focused taxonomy of propaganda techniques and compare it against an established, higher-agreement schema. |
| Outcome: | The proposed taxonomy outperforms existing models and reveals methodological differences hidden in base models. |
Copied to clipboard
| Challenge: | a typology is grounded in four linguistically motivated dimensions: form, evidentiality, epistemic stance, and tone. |
| Approach: | They propose a typology to evaluate how different EoBs affect whether models follow context versus prior knowledge. |
| Outcome: | The proposed model systematically evaluates 16 LLMs that differ in architecture, scale, and training stages . human listeners subconsciously interpret the belief based on how it is expressed, i.e., its explicitness, tone, or contextual cues. |
Copied to clipboard
| Challenge: | Existing data selection techniques are designed for small data pools, a study finds . filtering data by token length is an efficient method for improving results . |
| Approach: | They use self-scoring methods that do not rely on external help to perform fine-tuning . they also find that filtering data by token length offers a stable and efficient method . |
| Outcome: | The proposed methods outperform random selection on large datasets on large data pools. |
Copied to clipboard
| Challenge: | Existing methods for transferring knowledge from resource-rich domains to unknown domains are data hungry . a meta-learning algorithm is proposed to solve the problem of zero/few-shot DST . |
| Approach: | They propose a meta-learner for the problem of zero/few-shot DST . they propose to agnostically train any existing chatbot system to improve its performance . |
| Outcome: | The proposed meta-learner improves on baseline in a low-data setting. |
Copied to clipboard
| Challenge: | Large language models can generate questions with controlled difficulty, but they often fail to align with the given target difficulty. |
| Approach: | They propose a question generation method that requires no tuning of generator parameters yet significantly improves difficulty consistency. |
| Outcome: | The proposed method outperforms several mainstream methods on high-quality question answering datasets and achieves superior consistency with target difficulty. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated the effectiveness of self-alignment in which a large language model is aligned to follow general instructions using instructional data generated from the model itself. |
| Approach: | They propose to use human-written seeds to align large language models to follow general instructions to achieve cross-task generalization. |
| Outcome: | The proposed model outperforms base models and models that are generally instruction-tuned or have been adapted to the target domain by a large margin. |
Copied to clipboard
| Challenge: | Social networking services (SNS) have experienced rapid growth, which has proposed significant challenges for platform content management and interaction quality improvement. |
| Approach: | They propose a domain-specific LLM to break the performance bottleneck of single-task baselines and establish a comprehensive foundation for social networking services. |
| Outcome: | The proposed model achieves an average improvement of 14.02% across 8 major tasks and 7.56% in bilingual evaluation benchmark, compared with baseline models. |
Copied to clipboard
| Challenge: | Existing studies on suicide notes have not explored the topic of emotion detection. |
| Approach: | They develop a fine-grained emotion annotated corpus of suicide notes in English and use it to perform emotion detection on a curated dataset. |
| Outcome: | The proposed model performs emotion detection on a curated dataset of 205 suicide notes in English. |
Copied to clipboard
| Challenge: | Existing methods to optimize instruction-following capabilities of large language models (LLMs) assume that larger or stronger models are stronger teachers and therefore adopt smaller models as response generators. |
| Approach: | They propose to use large-scale instruction datasets to tune large language models to align with specific tasks and user intents. |
| Outcome: | The proposed metric outperforms most baselines in identifying the effectiveness of response generators. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are rapidly transforming the landscape of artificial intelligence due to the substantial resources required for training. |
| Approach: | They propose a post-deployment attack that bypasses system prompts to compromise models . they introduce Precise Activation Guarding and Unit Deviation Sampling to protect against attack . |
| Outcome: | The proposed attack bypasses system prompts, enabling unrestricted model outputs and safety violations. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are often evaluated using multiple-choice questions (MCQs) modeled on exams like the USMLE. |
| Approach: | They created a fictional medical benchmark centered on an imaginary organ, the Glianorex, to separate memorized knowledge from reasoning ability. |
| Outcome: | The proposed model outperforms base models in English but not in French. |
Copied to clipboard
| Challenge: | FLUKE introduces controlled variations across linguistic levels and leverages large language models with human validation to generate modifications. |
| Approach: | They propose a framework for assessing model robustness through systematic minimal variations of test data. |
| Outcome: | The proposed framework evaluates models and LLMs across six diverse NLP tasks and shows that they are more robust to natural, fluent modifications than base models. |
Copied to clipboard
| Challenge: | Large language models (LLMs) evolve to autonomous agents synthesizing real-time information, but their reasoning capabilities introduce an unexpected attack surface. |
| Approach: | They propose a framework that constructs deceptive narratives through adversarial debate and coordinated posting of evidence fragments, causing victims to internalize and propagate fabricated conclusions. |
| Outcome: | The proposed framework constructs deceptive narratives through adversarial debate and coordinated posting of evidence fragments, causing victims to internalize and propagate fabricated conclusions. |
Copied to clipboard
| Challenge: | generative Large Language Models (LLMs) are a promising tool for biomedical and healthcare research. |
| Approach: | They propose to use finetuned LLMs and multimodal LLM for genomic and proteomics tasks. |
| Outcome: | The proposed models outperform closed-source models in genomic and proteomics tasks and are highly accurate. |
Copied to clipboard
| Challenge: | Existing models claim to perform better on tasks measuring model capabilities, but there is no standard setup for reproducible evaluations. |
| Approach: | They propose a document that is documented and practical for reproducible LLM evaluations and includes recommendations from existing literature and new experiments. |
| Outcome: | The proposed standard identifies and reviews the varying factors in evaluation practices adopted by the community, such as prompt formatting, choice of in-context examples, probability normalizations, and task formulation. |
Copied to clipboard
| Challenge: | Sparse autoencoders (SAEs) have been proposed to mitigate polysemanticity, where neurons activate for multiple unrelated concepts. |
| Approach: | They propose a sparse autoencoder to transform dense activations into sparser, more interpretable features by transforming them into sparses. |
| Outcome: | The proposed model reduces polysemanticity and achieves higher concept separability. |
Copied to clipboard
| Challenge: | Existing methods to improve NAT model's performance but do not fully utilize it. |
| Approach: | They propose a non-autoregressive translation method which can obtain high-quality translations while maintaining the inference speed of NAT models. |
| Outcome: | The proposed method outperforms the autoregressive translation model on three translation tasks with 7.6 speedup. |
Copied to clipboard
| Challenge: | Despite the advances in large language models, they still face difficulties with multi-step reasoning tasks. |
| Approach: | They propose a method that randomly masks certain tokens within the chain of thought to improve model accuracy by 5% over standard supervised fine-tuning. |
| Outcome: | The proposed method improves accuracy and accuracy by 5% over standard fine-tuning with a few codes modified. |
Copied to clipboard
| Challenge: | Unsupervised parsing is a form of reinforcement learning that improves syntactic structures but lacks interpretability due to its lack of ad hoc heuristics. |
| Approach: | They propose an unsupervised approach that transfers syntactic knowledge to a Tree-LSTM model with discrete parsing actions. |
| Outcome: | The proposed model outperforms existing models on the All Natural Language Inference dataset and achieves a new state of the art in terms of parsing F-score. |
Copied to clipboard
| Challenge: | a large language model's (LLM) output distribution is changed by an alignment process . a recent study shows that aligned models surface information that cannot be recovered from base models without fine-tuning. |
| Approach: | They analyze two aspects of the alignment process that change output distributions . they find alignment suppresses irrelevant and unhelpful content . |
| Outcome: | The proposed model can be imitated without fine-tuning by using in-context examples and lower-resolution semantic hints about response content. |
Copied to clipboard
| Challenge: | Learning from human feedback has enabled the alignment of language models (LMs) with human preferences. |
| Approach: | They propose a Hybrid Preference routER that defers an annotation to either humans or LMs, achieving better annotation quality while reducing the cost of human-only annotation. |
| Outcome: | The proposed model achieves better annotation quality while reducing the cost of human-only annotation. |
Copied to clipboard
| Challenge: | Existing methods for related work generation (RWG) suffer from shallow comprehension due to taking the limited portions of references as input and isolated explanation for each reference due to ineffective capturing the relationships among them. |
| Approach: | They propose a multi-agent framework that takes the limited portions of references papers as input and isolates the relationships between them. |
| Outcome: | The proposed framework outperforms other selectors and improves reading order with constrains of the graph structure. |
Copied to clipboard
| Challenge: | Existing knowledge editing methodologies often encounter parameter conflict during knowledge overwriting and excessive computational overhead. |
| Approach: | They propose a method that erases outdated knowledge and inserts new knowledge at the location that corresponds to the target knowledge. |
| Outcome: | The proposed method achieves more effective knowledge editing at a lower cost compared to previous methods across various base models. |
Copied to clipboard
| Challenge: | Recent work probing pre-trained language models for downstream tasks is difficult to explain . a growing body of research is devoted to understanding what linguistic properties these language models have acquired. |
| Approach: | They propose a procedure and analysis method that takes a hypothesis of how a transformer-based model might encode a linguistic phenomenon and tests its validity. |
| Outcome: | The proposed method tests a hypothesis that some attention heads will consistently attend from a word in negation scope to the negation cue. |
Copied to clipboard
| Challenge: | Experimental results show that Prolog-MATH generates 81.3% solution coverage on Deepseek-V3 . |
| Approach: | They propose a curated corpus to support mathematical reasoning in large language models . they propose supervised fine-tuning followed by GRPO training to address problems that Deepseek-V3 fails to solve. |
| Outcome: | The proposed pipeline achieves 81.3% solution coverage on the Deepseek-V3 training set. |
Copied to clipboard
| Challenge: | Compared with LoRA and BitFit, training a single steering vector per layer with reinforcement learning requires orders of magnitude fewer resources and isolates a much smaller, more interpretable parameter set. |
| Approach: | They propose to train a single steering vector per layer with reinforcement learning while freezing all base weights to match the accuracy of fully RL-tuned reasoning models. |
| Outcome: | The proposed approach improves on an 8 billion-parameter model while keeping all base weights fixed. |
Copied to clipboard
| Challenge: | Experimental results demonstrate that our models achieve over 7% performance improvement compared to both SFT and RL-with-SFT models under the same experimental settings. |
| Approach: | They propose a dynamic generalization-guided reward design for rule-based RL that shifts rewards from exploratory to exploitative tool-use patterns. |
| Outcome: | The proposed model achieves over 7% performance improvement compared to SFT and RL-with-SFT models under the same experimental settings. |
Copied to clipboard
| Challenge: | Existing data selection methods suffer from severe domain specificity . existing methods for general instruction-following fail on reasoning tasks . |
| Approach: | They propose a framework that operationalizes contrastive entropy as a domain-adaptive selection criterion through warmup calibration, bi-directional NLL filtering, and entropic-based ranking. |
| Outcome: | Experiments show that InstructDiff outperforms baseline training on reasoning tasks while using only 10% of the data. |
Copied to clipboard
| Challenge: | Until 2021, most efforts were concentrated on one or two specific tasks such as error detection (ED) and data imputation (DI). |
| Approach: | They propose to instruction tune local LLMs as universal DP task solvers that operate on a local, single, and low-priced GPU, ensuring data security and enabling further customization. |
| Outcome: | The proposed models deliver competitiveness and generalizability to unseen tasks while barely compromising the base models’ abilities in NLP tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly deployed in security-sensitive applications . recent defenses rely on supervised fine-tuning with benign and malicious labels . position bias arises when benign content placed later in a prompt is rejected at much higher rates . |
| Approach: | They analyze three recurring shortcut behaviors induced by supervised fine-tuning . position bias arises when benign content placed later in a prompt is rejected . token trigger bias occurs when strings common in attack data raise rejection probability . |
| Outcome: | The proposed model overrides intended logic when adversarial instructions appear . the proposed model has low rejection rates but narrow correlations in defense data . |
Copied to clipboard
| Challenge: | Existing methods for evaluating CNs are expensive, time-consuming, and subjective, but lack a universal truth and the lack of a 'universal truth' . |
| Approach: | They propose a model ranking pipeline based on pairwise comparisons of generated CNs from different models organized in a tournament-style format to improve the evaluation process. |
| Outcome: | The proposed method achieves a high correlation with human preference, with a score of 0.88, and compares chat, instruct, and base models, exploring their strengths and limitations. |
Copied to clipboard
| Challenge: | Low-Rank Adaptation (LoRA) is one of the most efficient parameter-efficient fine-tuning methods. |
| Approach: | They propose to conceptualize each LoRA module as a beam where each rank corresponds to a potential sub-solution. |
| Outcome: | The proposed method improves performance on three base models and 12 datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) require alignment to effectively and safely follow user instructions. |
| Approach: | They propose a simple, training-free algorithm that aligns any base model at inference time using a small aligned model. |
| Outcome: | The proposed algorithm outperforms large aligned models on open-instruction tasks without training. |
Copied to clipboard
| Challenge: | Low-resource methods for LLM alignment have been popular, but still face challenges in obtaining high-quality and aligned content. |
| Approach: | They propose a framework to enhance alignment ability of base models by the guidance of a small aligned model. |
| Outcome: | The proposed framework outperforms baseline methods while avoiding degradation on downstream tasks. |
Copied to clipboard
| Challenge: | Existing methods for inference are often myopic and have divergent reasoning paths . a meta-adaptive reasoning framework is proposed to improve the efficiency of LLM agents . |
| Approach: | They propose a meta-adaptive reasoning framework that integrates tool execution and reasoning planning. |
| Outcome: | The proposed framework outperforms existing methods in performance and inference efficiency. |
Copied to clipboard
| Challenge: | Several studies claim that domain-adaptive pretraining improves performance on downstream medical tasks. |
| Approach: | They compare medical LLMs and VLMs against their corresponding base models . they find that medical Lms outperform their base models in 12.1% of cases . |
| Outcome: | The proposed models outperform their base models on medical questions and tasks in 12.1% of cases and reach a tie in 49.8% of cases. |
Copied to clipboard
| Challenge: | Existing benchmarks for code generation tasks are inadequate, but performance declines on self-invoking tasks. |
| Approach: | They propose a general recipe for generating more challenging versions of existing benchmarks . they propose to use instruction-tuned models to evaluate LLMs on self-invoking code generation tasks . |
| Outcome: | The proposed model improves on humanEval and MBPP but on self-invoking code generation tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) still exhibit significant deficiencies in basic language understanding and manipulation. |
| Approach: | They propose a bilingual benchmark to assess the performance of Large language models . they use a set of 15 simple text editing tasks to examine their capabilities . |
| Outcome: | The proposed benchmark aims to assess the performance of Large language models in basic language tasks. |
Copied to clipboard
| Challenge: | Existing solutions to control speaker-related gender inflections in ST involve dedicated model retraining on gender-labeled data. |
| Approach: | They propose to use a gender-based inference-time solution to control speaker-related gender inflections in ST by replacing the implicitly learned internal language model with gender-specific external LMs. |
| Outcome: | The proposed approach outperforms the base models and the best training-time mitigation strategy by up to 31.0 and 1.6 points in gender accuracy, respectively, for feminine forms. |
Copied to clipboard
| Challenge: | Reward Informed Fine-Tuning (RIFT) is an effective and robust alternative to expensive expert data for LLM alignment. |
| Approach: | They propose a reward-informed fine-tuning framework that utilizes all self-generated samples to learn from both positive and negative trajectories. |
| Outcome: | The proposed framework outperforms both RFT and Supervised Fine-Tuning (SFT) on mathematical benchmarks. |
Copied to clipboard
| Challenge: | Multi-agent debate is compute-intensive and requires long transcripts before answering questions. |
| Approach: | They propose a framework that distills multi-agent debate into a single LLM by combining debate structure learning with internalization via dynamic reward scheduling and length clipping. |
| Outcome: | The proposed model matches or exceeds explicit multi-agent debate performance using 93% fewer tokens across multiple models and benchmarks. |
Copied to clipboard
| Challenge: | Parameter-efficient fine-tuning (PEFT) is a common method for fine- tuning large language models . however, once updated, PEFT modules suffer performance degradation on newer versions . |
| Approach: | They propose a method that enhances the PEFT module by focusing on the task-specific pattern while reducing its dependence on certain knowledge in the base model. |
| Outcome: | Experiments show that PEFT modules can maintain performance on updated models without re-tuning . the proposed approach can be used in real-world applications with large model sizes . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable capabilities in Tool-Integrated Reasoning (TIR) however, the practical application is often hindered by frequent errors in tool invocations, such as incorrect tool names, invalid parameters, wrong tool-call order, or malformed invocation formats. |
| Approach: | They propose a specialized post-processing module that performs independent reasoning on the input of a frozen upstream LLM and an advanced RL algorithm to improve the tool-use reliability of base LLMs. |
| Outcome: | The proposed module improves task completion rates and invocation accuracy over the raw outputs of various upstream LLMs on a diverse set of tool-use and reasoning benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have impressive capabilities in comprehending human language and vast parametric knowledge obtained from large corpora. |
| Approach: | They propose a multi-level benchmark for free text model editing to bridge the gap . they categorize probe queries into three levels of generalization . |
| Outcome: | The proposed method improves the generalization performance of large langugae models. |
Copied to clipboard
| Challenge: | a long-standing debate concerns whether the linguistic input children receive is sufficient to explain the grammatical knowledge they develop. |
| Approach: | They evaluate baby language models trained on child-oriented input from the BabyLM Challenge and two base models trained in 10M and 100M tokens. |
| Outcome: | The proposed models acquire filler-gap dependencies but fail to generalize or fully capture island constraints. |
Copied to clipboard
| Challenge: | Existing evaluation methods for large language models (LLMs) are inadequate to provide solid conclusions for key experiments such as data ablation and scaling law. |
| Approach: | They propose a method specifically designed to optimize the evaluation of base models by incorporating two innovations: In-Context Light-instruction Prompt and Blank-ppl for multi-choice tasks with candidate options. |
| Outcome: | The proposed method significantly improves stability and consistency of evaluations during pre-training and consistency between base and instruct models. |
Copied to clipboard
| Challenge: | We empirically show that Heaps’ and Zipf’s laws only hold for LLM-generated texts in a narrow model-dependent temperature range. |
| Approach: | They propose to apply Zipf's law to large language models to study the frequency distribution of words in human-written texts . |
| Outcome: | The proposed models only hold for LLM-generated texts in a narrow model-dependent temperature range. |
Copied to clipboard
| Challenge: | Large language models have shown promise in clinical decision making, but current approaches struggle to localize and correct reasoning errors at specific steps of the reasoning process. |
| Approach: | They propose a process reward modeling framework that leverages retrieval-augmented generation to verify each reasoning step against established medical knowledge bases. |
| Outcome: | The proposed model improves on five medical QA benchmarks and two open-ended diagnostic tasks by 13.50% on MedQA. |
Copied to clipboard
| Challenge: | Recent self-training approaches have reduced reliance on human-labeled data, which limits their scalability. |
| Approach: | They propose a team-based self-play algorithm that iteratively refines alignment without additional human supervision. |
| Outcome: | The proposed algorithm outperforms baselines and LLM benchmarks in the self-supervised setting. |
Copied to clipboard
| Challenge: | Recent work has focused on improving the mathematical reasoning capabilities of Large Language Models (LLMs). |
| Approach: | They propose an end-to-end framework to integrate FL into NL math reasoning . they propose a problem alignment method that reformulates QA and existence problems . |
| Outcome: | The proposed framework achieves 89.80% and 84.34% accuracy rates on the MATH-500 and the AMC benchmarks. |
Copied to clipboard
| Challenge: | Inference-time alignment approaches still face limitations due to policy-specific value functions and latency during the inference phase. |
| Approach: | They propose an efficient and policy-agnostic preference optimization method that avoids time latency associated with token generation. |
| Outcome: | The proposed method achieves a favorable trade-off between alignment quality and inference-time latency. |
Copied to clipboard
| Challenge: | Existing reward models lack generative and reasoning capabilities, resulting in poor performance. |
| Approach: | They propose a reward-aware task-adaptive reward model that enables pointwise training using readily available pairwise data via a novel Preference-Aware Reward mechanism. |
| Outcome: | The proposed reward model achieves an average relative improvement of 8.7% over the base models on RewardBench and RMBench. |
Copied to clipboard
| Challenge: | Large language models pre-trained on massive data have promoted multilingual natural language processing (NLP). |
| Approach: | They construct a bilingual translation corpus with 2,500 language pairs and develop a suite of four models with parallel data. |
| Outcome: | The proposed model suites are evaluated across 7 tasks and 12 benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) reach hundreds of billions of parameters and require resources for training and inference stages. |
| Approach: | They propose a low-rank adapter to reduce the number of trainable parameters in a model and reduce memory requirements. |
| Outcome: | The proposed approach reduces memory and compute requirements while preserving performance. |
Copied to clipboard
| Challenge: | Speculative decoding is a prominent technique for accelerating LLM inference by leveraging an auxiliary draft model, but its effectiveness is limited by the autoregressive nature of draft generation. |
| Approach: | They propose a method that integrates speculative draft generation directly within the target model using multi-stream attention. |
| Outcome: | The proposed method improves acceptance but also latency and speculation latency, limiting overall speedup. |
Copied to clipboard
| Challenge: | Nguni languages have over 20 million home language speakers in South Africa . there has been considerable growth in the datasets for these languages, but no analysis of the performance of NLP models for these language has been reported across languages and tasks. |
| Approach: | They compile publicly available datasets for natural language understanding and generation, spanning 6 tasks and 11 datasets. |
| Outcome: | The proposed models outperform existing models and large-scale adapted models on cross-lingual transfer and machine translation. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been used to remove harmful knowledge and undesirable capabilities. |
| Approach: | They propose a framework that leverages Cognitive Diagnosis Modeling to evaluate LLM unlearning. |
| Outcome: | The proposed framework enhances evaluation and facilitates removal of harmful abilities. |
Copied to clipboard
| Challenge: | Existing approaches to incentivize LLMs’ deep thinking abilities require large-scale data or significant training efforts. |
| Approach: | They introduce an efficient framework that enhances LLM reasoning by teaching models to self-verify and self-correct during inference. |
| Outcome: | The proposed framework outperforms models trained on long-CoT distilled data with 3.1k initialization samples and achieves an accuracy improvement of 51.0% to 81.6%. |
Copied to clipboard
| Challenge: | Existing methods for process supervision fail to distinguish meaningful progress from mere verbosity . existing methods lack a coherent approach to process supervision . |
| Approach: | They propose a framework that formalizes reasoning as a trajectory through a state space of empirical solvability. |
| Outcome: | The proposed framework achieves an average accuracy gain of 3% with 30% reduced token consumption. |
Copied to clipboard
| Challenge: | Existing studies have overlooked the impact of hyperparameters on table understanding abilities . authors show that smaller learning rates and fewer training instances can enhance table understanding while preserving general capabilities. |
| Approach: | They propose a hyperparameter-based instruction-tuned model for table-related tasks that improves out-of-domain table understanding ability and general capabilities. |
| Outcome: | The proposed model outperforms existing models on table-related tasks while maintaining strong out-of-domain generalization and general capabilities. |
Copied to clipboard
| Challenge: | Small Language Models (SLMs) are becoming increasingly popular in specialized fields such as industrial applications. |
| Approach: | They propose a framework which transfers reasoning capabilities via Chain-of-Thought distillation from Large Language Models (LLMs) to smaller, more efficient models (SLMs) |
| Outcome: | The proposed framework outperforms the base models in Industry 4.0 by a significant margin. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have greatly improved natural language understanding and generation. |
| Approach: | They train a wide range of base models on a variety of datasets including code generation, mathematical reasoning, and general-domain tasks. |
| Outcome: | The results show that training–task synergies persist across all models while others vary substantially, emphasizing the importance of model-specific strategies. |
Copied to clipboard
| Challenge: | Pre-trained language models have advanced natural language processing (NLP) despite the introduction of BERT, single-language models are still relevant. |
| Approach: | They present a German singlelanguage RoBERT model pre-trained exclusively on the German portion of the OSCAR dataset. |
| Outcome: | The GottBERT model outperforms the existing models on Named Entity Recognition and text classification tasks. |
Copied to clipboard
| Challenge: | Empirical evaluations on eight recent LLMs reveal that DRPO significantly enhances alignment performance, enabling base models to outperform their SFT/RLHF-tuned counterparts. |
| Approach: | They propose a tuning-free approach to self-alignment called Dynamic Rewarding with Prompt Optimization (DRPO) it leverages a dynamic rewarding mechanism to identify and rectify alignment weaknesses . |
| Outcome: | The proposed approach outperforms existing methods and is highly adaptable to various alignment challenges. |
Copied to clipboard
| Challenge: | Existing models generate tokens by updating high-dimensional representations and decoding from them at each timestep. |
| Approach: | They propose a framework that allows reasoning correction and length control based on derived ideal trajectories. |
| Outcome: | The proposed model can predict correctness and length control based on ideal trajectories. |
Copied to clipboard
| Challenge: | Pretrained LLMs fail to capture behavioral diversity of target populations due to inherent variability across individuals and groups. |
| Approach: | They propose a probabilistic prompting method that aligns LLM responses with the target population. |
| Outcome: | Experiments show that the proposed method outperforms competing methods in alignment and diversity metrics. |
Copied to clipboard
| Challenge: | Existing studies highlight that large language models are receptive to external information that contradicts their parametric knowledge, but little research has been conducted on the direct impact of instruction-tuning on this phenomenon. |
| Approach: | They examine how instruction-tuning influences LLMs' susceptibility to misinformation, particularly in knowledge conflict situations. |
| Outcome: | The proposed model is more user-oriented and more likely to accept misinformation when it is presented by the user. |
Copied to clipboard
| Challenge: | 3GPP standards define the technical design and implementation of 5G systems . expert-level questions require navigating thousands of pages of cross-referenced standards . |
| Approach: | They propose a standard-native retrieval-augmented generation system that can answer 5G questions . they use SpecDB, ChangeDB, TDocDB and a metadata-rich retrieval system to do this . |
| Outcome: | The proposed solution outperforms base models and state-of-the-art RAG systems in QA datasets . expert-level queries require navigating thousands of pages of cross-referenced standards . |
Copied to clipboard
| Challenge: | Multi-document (MD) processing is crucial for LLMs to handle real-world tasks such as summarization and question-answering across large sets of documents. |
| Approach: | They propose a framework that generates high-quality synthetic MD instruction data over sets of articles via targeted prompts. |
| Outcome: | MDCure generates high-quality synthetic MD instruction data over sets of articles . evaluations show it improves over pre-trained models by up to 75.1% . |
Copied to clipboard
| Challenge: | Existing methods for self-improvement of large language models with verifiable rewards (RLVR) can drift over iterations, while corpus-grounded approaches rely on curated data environments. |
| Approach: | They propose a Web-grounded Iterative Self-play Tree framework for domain-targeted reasoning improvement that learns directly from the open-web without requiring any pre-arranged domain corpus. |
| Outcome: | The proposed framework outperforms both purely endogenous self-evolution and corpus-grounded self-play baselines and is domain-steerable. |
Copied to clipboard
| Challenge: | Existing approaches for SPARQL generation rely on one-turn models. |
| Approach: | They propose a training-free interactive refinement pipeline that acts as a plug-and-play enhancement for existing SPARQL systems. |
| Outcome: | The proposed approach improves the accuracy of base models without fine-tuning . it transforms potentially flawed queries from any source into verifiable code . |
Copied to clipboard
| Challenge: | Reinforcement learning (RL) has emerged as a powerful paradigm for improving the reasoning capabilities of large language models. |
| Approach: | They propose a pipeline that automatically discovers thinking token patterns with reasoning primitives and curates SFT datasets to prepare LLMs for RL. |
| Outcome: | The proposed pipeline outperforms baseline methods on mathematical and logical reasoning benchmarks on RL tasks. |
Copied to clipboard
| Challenge: | Existing Chain-of-Thought (CoT) methods struggle with consistency and verification in complex reasoning tasks. |
| Approach: | They propose a framework that integrates structured knowledge representation with learned planning. |
| Outcome: | The proposed framework outperforms existing Chain-of-Thought (CoT) methods on math reasoning, logical reasoning, and coding tasks. |
Copied to clipboard
| Challenge: | Recent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning. |
| Approach: | They propose a GRPO-based interactive training approach that augments the rollouts of a student model with the guidance of . a teacher poses a problem, lets the student try an answer, then gives corrective feedback–enough to point the mind in the right direction and then show the correct solution. |
| Outcome: | The proposed method shows 3.69% improvement over zero-shot baselines and 2.08% and 3.99% improvement over the vanilla-GRPO baselines. |
Copied to clipboard
| Challenge: | Existing methods for reasoning-intensive information retrieval suffer from inefficiency . Chain-of-Thought (CoT) approaches suffer from lack of token efficiency . Existing models lack episodic memory, which stores the history of prior states . |
| Approach: | They propose an algorithm that enhances state-based frameworks with an episodic memory module that stores the full history of prior states for a query. |
| Outcome: | The proposed model outperforms CoT and state-based models on the BRIGHT benchmark and is highly token-efficient. |
Copied to clipboard
| Challenge: | State-of-the-art multimodal language models (MLMs) show promise for supporting SLPs, but their use remains underexplored due to a limited understanding of their performance in high-stakes clinical settings. |
| Approach: | They propose a taxonomy of real-world use cases of multimodal language models in speech-language pathologies to address this gap. |
| Outcome: | The proposed model outperforms 15 state-of-the-art models in speech-language pathologies across five use cases and achieves improvements of over 30% on domain-specific data. |
Copied to clipboard
| Challenge: | *SchED* is a training-free, model-agnostic early-exit algorithm that terminates diffusion decoding using a progress-aware confidence threshold. |
| Approach: | They propose a training-free, model-agnostic early-exit algorithm that terminates diffusion decoding using a progress-aware confidence threshold. |
| Outcome: | The proposed algorithm achieves 4 speedups on instruction-tuned models while maintaining baseline performance on average. |
Copied to clipboard
| Challenge: | Experiments show that enhancing implicit reasoning capabilities can significantly improve complex instruction following in large language models. |
| Approach: | They propose a method to enhance LLMs’ understanding of implicit reasoning instructions by formalizing such instructions as verifiable reasoning graphs and fine-tuning with graph reasoning. |
| Outcome: | The proposed method outperforms existing models on five complex instruction following benchmarks and will be open-sourced in the near future. |
Copied to clipboard
| Challenge: | **ViLegalLM** is the first suite of Vietnamese pretrained language models for legal text understanding and generation. |
| Approach: | They propose a suite of Vietnamese pretrained language models for legal text understanding and generation. |
| Outcome: | The proposed models outperform instruction-tuned adaptation on four main Vietnamese legal downstream tasks. |
Copied to clipboard
| Challenge: | Existing iterative refinement strategies that generate solutions in a single forward pass often hit a performance ceiling on complex algorithmic tasks. |
| Approach: | They propose a reinforcement learning framework that internalizes the structured reasoning trajectory directly into the model’s weights. |
| Outcome: | The proposed framework achieves 94.51% (87.20%) on HumanEval, 81.80% (78.57%) on MBPP, 35.00% on BigCodeBench, 52.21% on LiveCodeBech, and 37.34% on CodeForces in a single-attempt setting. |
Copied to clipboard
| Challenge: | Preference alignment methods can reinforce hallucinations when preference judgments reward fluency and confidence over factual correctness. |
| Approach: | They propose a method that corrects misordered preference pairs and adds a factuality-aware margin to emphasize pairs with clear correctness differences. |
| Outcome: | The proposed method improves factuality and reduces hallucination rates across seven open-weight LLMs. |
Copied to clipboard
| Challenge: | Multi-agent LLMs are rapidly moving from prototype to real-world use . network topology is a first-order security parameter in multi-aggent systems . |
| Approach: | They propose a framework for comparing topology-conditioned memory leakage in multi-agent LLM systems. |
| Outcome: | The proposed framework evaluates topology-conditioned memory leakage in multi-agent LLM systems. |
Copied to clipboard
| Challenge: | Existing user simulators based on prompting to role-play or SFT focus on imitating textual utterances without considering multi-faceted cognitive processes that underlie human decision-making during interactions. |
| Approach: | They construct a user-simulator dataset that augments 51k human–LLM conversations by reconstructing the user’s inner reasoning during and at the end of each dialogue. |
| Outcome: | The proposed user simulators augment 51k human–LLM conversations by reconstructing the user’s inner reasoning both during and at the end of each dialogue. |